A Generalized Reinforcement-Learning Model: Convergence and Applications
نویسندگان
چکیده
Reinforcement learning is the process by which an autonomous agent uses its experi ence interacting with an environment to im prove it::; beha.vior. The rVlarkov dec.ision pro cess (:\'lDP) modd is a popular way of for malizing the reinforcement-learning problem, but it it) by no meant; LIte only \-vay. In Lhitl paper, we show hmy IllaIlY of the important theoretical results concerning reinforcement learning in MDPs extend to a generalized MDP model that includes MDPS, two-player games and .\lUPs under a worst-case optimality cri terion as special cases. The basis of this ex tension is a stochastic-approximation theo rem that redllces fLsynchronolls convergence to synchrOflOlJS cOlnrergence.
منابع مشابه
A Generalized Reinforcement-Learning Model: Convergence and Applicationa
Reinforcement learning is the process by which an autonomous agent uses its experience interacting with an environment to improve its behavior. The Markov decision process (mdp) model is a popular way of formalizing the reinforcement-learning problem, but it is by no means the only way. In this paper, we show how many of the important theoretical results concerning reinforcement learning in mdp...
متن کاملOperation Scheduling of MGs Based on Deep Reinforcement Learning Algorithm
: In this paper, the operation scheduling of Microgrids (MGs), including Distributed Energy Resources (DERs) and Energy Storage Systems (ESSs), is proposed using a Deep Reinforcement Learning (DRL) based approach. Due to the dynamic characteristic of the problem, it firstly is formulated as a Markov Decision Process (MDP). Next, Deep Deterministic Policy Gradient (DDPG) algorithm is presented t...
متن کاملReinforcement Learning in Neural Networks: A Survey
In recent years, researches on reinforcement learning (RL) have focused on bridging the gap between adaptive optimal control and bio-inspired learning techniques. Neural network reinforcement learning (NNRL) is among the most popular algorithms in the RL framework. The advantage of using neural networks enables the RL to search for optimal policies more efficiently in several real-life applicat...
متن کاملReinforcement Learning in Neural Networks: A Survey
In recent years, researches on reinforcement learning (RL) have focused on bridging the gap between adaptive optimal control and bio-inspired learning techniques. Neural network reinforcement learning (NNRL) is among the most popular algorithms in the RL framework. The advantage of using neural networks enables the RL to search for optimal policies more efficiently in several real-life applicat...
متن کاملGeNGA: A Generalization of Natural Gradient Ascent with Positive and Negative Convergence Results
Natural gradient ascent (NGA) is a popular optimization method that uses a positive definite metric tensor. In many applications the metric tensor is only guaranteed to be positive semidefinite (e.g., when using the Fisher information matrix as the metric tensor), in which case NGA is not applicable. In our first contribution, we derive generalized natural gradient ascent (GeNGA), a generalizat...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 1996